Back

Pharmacoepidemiology and Drug Safety

Wiley

Preprints posted in the last 30 days, ranked by how well they match Pharmacoepidemiology and Drug Safety's content profile, based on 18 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Systematic Data Fitness Assessment Improves Validity and Replicability of Research Using Real-World Data

Razzaghi, H.; Wieand, K.; Pinkney, A.; Bailey, C.

2026-08-10 epidemiology 10.64898/2026.08.05.26359818 medRxiv
Top 0.1%
10.1%
Show abstract

Research replication is essential to build trust in evidence produced from real-world data. However, methods for conducting and reporting these studies are lacking, particularly related to data quality and fitness assessments. We replicated a single-center study from Children's Hospital of Atlanta in a multi-institutional learning network (PEDSnet) to evaluate the long-term effects of hydroxyurea in children with severe sickle cell disease (SS/S{beta}0 genotype). An AS-IS arm applied the original study's criteria with no major data quality adjustments, while a Data Fitness Enhanced (DFE) arm used systematic data fitness assessment to inform adjustments to cohort inclusion criteria and variable definitions; both arms then replicated the original study's primary analyses. Data quality checks in the DFE arm refined cohort criteria and improved hydroxyurea capture, drug era computation, and hematology specialist mapping. The DFE cohort produced average treatment effects with higher face validity and greater concordance with the original study (e.g., change in ED visits: -0.44 (CI -0.60, -0.26) versus -0.36 (CI -0.57, -0.16) in the original study) than the AS-IS cohort (-0.08 (CI -0.26, 0.09)), which yielded several implausible results. These findings show that superficially plausible cohort characteristics do not guarantee valid results without transparent, systematic data fitness assessment.

2
A Renal Safety Checkpoint for Early High Intensity Statin Therapy in Critically Ill Patients With Acute Coronary Syndrome: A Multidatabase Target Trial Emulation

Huang, K.; Zheng, X.; Liu, J.; Wu, C.; Sun, H.

2026-08-07 cardiovascular medicine 10.64898/2026.08.05.26359828 medRxiv
Top 0.1%
6.6%
Show abstract

Background: High intensity statins are foundational after acute coronary syndrome (ACS), yet intensive care unit prescribing occurs while renal reserve, perfusion, and interacting therapies are changing. We tested a renal safety checkpoint integrating kidney status, hemodynamic instability, and drug interaction burden to identify when statin intensity may become nonexchangeable. Methods: We emulated an active-comparator target trial across MIMIC-IV, eICU, and MIMIC-III. Critically ill adults with ACS, acute myocardial infarction, or percutaneous coronary intervention who received high- or moderate-intensity statins within 24 hours were included. The primary outcome was 7-day KDIGO stage 2 or 3 acute kidney injury or incident renal replacement therapy. Eligibility, time zero, treatment assignment, and follow-up were aligned. Database-specific propensity scores, overlap weighting, and standardization addressed confounding and treatment overlap. Safety domains, longitudinal analyses, bootstrap resampling, source omission, and endpoint sensitivities assessed robustness. Results: Among 5,178 patients, 761 developed the primary outcome, including 223 who initiated renal replacement therapy. Standardized risks were 17.40% with high-intensity therapy and 15.01% with moderate-intensity therapy (risk difference, 2.39 percentage points [95% confidence interval (CI), -0.23 to 5.05]; risk ratio, 1.16 [95% CI, 0.99 to 1.39]). Risk separation was greatest with high hemodynamic instability (5.78 percentage points [95% CI, 1.56 to 9.74]) and high drug-interaction burden (6.24 percentage points [95% CI, -0.44 to 12.19]). Renal replacement therapy showed a 1.33-point risk difference (95% CI, 0.18 to 2.67). Conclusions: This study moves statin safety assessment beyond fixed dose label or isolated creatinine measurement. The findings support a clinically actionable monitoring strategy in which early statin intensity is reassessed against evolving perfusion, kidney status, and interaction burden. This approach preserves intensive lipid lowering for physiologically suitable patients while identifying a high risk window in which temporary moderation.

3
Publication Bias in Abstracts Presented at the American Diabetes Association Scientific Sessions: A Retrospective Cohort Study

Pinedo-Torres, I.; Taype-Rondan, A.; Vera-Luza, A. A.; Zegarra-Lizana, P. A.; Rojas-Vilca, J. L.; Yovera-Aldana, M.

2026-08-31 epidemiology 10.64898/2026.08.26.26361486 medRxiv
Top 0.1%
4.2%
Show abstract

Objective. To determine the publication rate of abstracts presented at the American Diabetes Association Scientific Sessions and to evaluate the association between statistical significance of study results and subsequent publication. Research Design and Methods. We conducted a retrospective cohort study of abstracts presented at the 2018 American Diabetes Association Scientific Sessions. The primary exposure was study result category (statistically significant vs. non-statistically significant findings), and the primary outcome was publication in an indexed journal within 5 years after conference presentation. Publication status was determined through PubMed/MEDLINE and Scopus searches. Adjusted relative risks (RRs) and 95% CIs were estimated using generalized linear models with Poisson distribution and robust variance. Results. Among 541 included abstracts, 321 (59.3%) were subsequently published in indexed journals. Abstracts reporting statistically significant findings had a higher publication rate than those reporting non-statistically significant findings (61.9% vs. 42.3%; p=0.002). In the adjusted analysis, abstracts with non-statistically significant findings had a lower likelihood of publication compared with those reporting statistically significant findings (adjusted RR 0.71 [95% CI 0.55-0.93]; p=0.013). Conclusions. Approximately four in ten abstracts presented at the ADA Scientific Sessions were not published within 5 years. Abstracts reporting non-statistically significant findings had a lower likelihood of subsequent publication, suggesting persistent publication bias in diabetology research. Future initiatives promoting the interpretation of effect estimates, confidence intervals and clinical relevance, rather than statistical significance alone, may help reduce selective dissemination of evidence

4
A Simulation Study Comparing Multiple Imputation and Complete Case Analysis for Handling Missing Preschool Body Mass Index

Savu, A.; Dover, D. C.; Hajihosseini, M.; Gaudet, L. A.; Kaul, P.

2026-08-14 epidemiology 10.64898/2026.08.13.26360115 medRxiv
Top 0.1%
4.1%
Show abstract

Background and Objective. Missing data frequently occurs in health databases and can bias analyses if not correctly dealt with. Using real-world data, we compared complete-case and multiple-imputation methods for recovering true parameters of a multivariable logistic regression model for the association between maternal glucose levels during pregnancy and child excess weight at preschool age, where missing values were present in as much as 30% of our sample. Methods. This study utilized a cohort of 130,424 children with complete preschool-age body mass index (BMI) measurements from the Calgary and Edmonton health regions of Alberta, Canada. In the complete BMI data, we introduced missingness through deletion following three distinct mechanisms: missing completely at random (MCAR), at random (MAR), and not at random (MNAR). To handle the missing data created, we employed complete-case and multiple-imputation methods. Maternal glucose levels during pregnancy were categorized into five groups and its association with child excess weight at pre-school age was determined based on a logistic regression model using the full observed data (yielding true values), observed data that was not deleted (complete-case estimates), and imputed data (multiple-imputation estimates). The accuracy of complete-case and multiple-imputation estimates were evaluated against the true values. Finally, we conducted a sensitivity analysis for the MNAR mechanism using pattern-mixture models with an additive shift. Results. Under MCAR and MAR, multiple-imputation generally outperformed complete-case, yielding smaller absolute and relative bias. Both methods achieved high significance ([≥] 0.96) for most effects. Mean squared errors for multiple-imputation and complete-case were similar missing completely at random, missing at random, and coverage was consistently high ([≥] 0.99). Under MNAR, both complete-case and multiple-imputation showed poor performance regarding bias and statistical significance. Sensitivity analysis using pattern-mixture models indicated performance varied by specific effect. Conclusions. Under MCAR and MAR, multiple-imputation introduced higher bias but demonstrated superior overall performance based on mean squared error and restored statistical power. Conversely, both methods failed under MNAR, where pattern-mixture modeling sensitivity analyses revealed highly variable, effect-specific performance due to unverifiable shift assumptions. When faced with missing data, researchers should assess missingness mechanisms, report both complete-case and multiple-imputation estimates under MCAR/MAR while accounting for power-versus-bias tradeoffs, and employ pattern-mixture sensitivity analyses to test robustness when MNAR is plausible.

5
CHARMS and PROBAST+AI: an updated template for Data Extraction and Risk of Bias Assessment in systematic reviews of prediction models

Jaber, A.; Hughes, L.; Cameron, A. C.; Quinn, T. J.

2026-08-31 cardiovascular medicine 10.64898/2026.08.26.26361189 medRxiv
Top 0.1%
3.2%
Show abstract

Background: Systematic reviews of clinical prediction models increasingly include studies using artificial intelligence (AI) and machine learning (ML) methods alongside traditional multivariable regression approaches. A previously published Excel tool enabled standardised data extraction using the CHARMS checklist and risk of bias assessment using PROBAST. The recent publication of the PROBAST+AI framework, which distinguishes the assessment of model development quality from the assessment of model evaluation risk of bias and assesses applicability in both parts, necessitates an updated digital instrument applicable across prediction modelling methods. Methods: We updated an open-access Excel tool to incorporate the full PROBAST+AI framework. The updated template incorporates structural separation between assessment of model development quality and model evaluation risk of bias, with applicability assessed in both parts. It also incorporates updated signalling questions, including those addressing methodological issues particularly relevant to AI/ML, and automates the generation of summary tables and graphical displays. Results: The updated tool (CHARMS & PROBAST+AI Template) contains 11 worksheets and supports data extraction and appraisal for up to 30 prediction models. Dedicated, linked worksheets enable separate assessment of model development and model evaluation, with Domain 4 distinguishing among Apparent, Internal, and External evaluation settings. Key updates include dedicated assessments for predictor pre-processing, class imbalance handling and recalibration, data leakage prevention, and replication of the full model development pipeline within resampling procedures. Automated sheets dynamically format tables and summary charts covering PROBAST+AI parts. Conclusions: The CHARMS & PROBAST+AI Excel template provides a standardised, user-friendly, and rigorous digital framework for systematic reviewers appraising traditional statistical and AI-driven clinical prediction models.

6
Twelve-Year Real-World Evaluation of a Regulated Guideline-Based Warfarin Dosing and Care Automation System

Tiihonen, M.

2026-08-12 health informatics 10.64898/2026.08.10.26360059 medRxiv
Top 0.1%
2.7%
Show abstract

Background: Warfarin therapy requires repetitive dose adjustments based on INR (International Normalised Ratio) monitoring. We evaluated the long-term real-world performance of Forsante Warfarin Advisor (WA), a CE-marked class IIb guideline-based decision support and care automation medical device used in anticoagulation management. Methods: Retrospective real-world data from routine clinical use between 2016 and 2026 were analysed. Treatment quality was assessed using Time in Therapeutic Range (TTR). Recommendation performance was evaluated by comparing achievement of target INR after clinician acceptance or modification of Warfarin Advisor recommendations. Results: Among 1348 patients in March 2026 median TTR was 83%, compared with 70% in March 2016. Dosages congruent with Warfarin Advisor recommendations were strongly associated with achieving target INR at follow-up in INR target ranges of 2.0-3.0 and 2.5-3.5. Treatment quality remained consistently high across years of deployment. No serious device-attributable safety incidents, regulatory incident reports, or CAPA cases were identified during 12 calendar years and 82,709 patient years of routine use. Conclusions: The findings provide real-world long-term evidence that a guideline-based warfarin dosing and care automation system can support sustained high-quality anticoagulation control in routine clinical practice. The findings support the feasibility of deploying workflow-integrated execution of selected guideline-driven clinical processes, while the causal effects on clinical outcomes require prospective confirmation. Keywords: Clinical decision support systems, Guideline execution, Real-world evidence, Warfarin, Anticoagulation

7
Beyond Signal Detection: Sequential Target Trial Emulations to Confirm Previously Detected Adverse Drug Event Signals for Atorvastatin in Older Medicare Beneficiaries

Rowan, C. G.

2026-08-14 epidemiology 10.64898/2026.08.12.26360302 medRxiv
Top 0.2%
2.6%
Show abstract

Importance: Active pharmacovigilance via sequential target trial emulation can detect adverse drug event (ADE) signals missed by spontaneous reporting, yet signals identified through high-dimensional screening require rigorous, pre-specified confirmation that addresses residual confounding, outcome heterogeneity, multiplicity, and absolute risk. Objective: To confirm or refute previously detected ADE signals associated with atorvastatin initiation among older adults by applying refined and more homogeneous outcome definitions, expanded family- and component-level exclusions, within-outcome false-discovery-rate control, and probabilistic quantitative bias analysis within a sequential target-trial framework. Design, Setting, and Participants: Confirmatory sequential target trial emulation study using Medicare fee-for-service claims (2017-2019). Eligible participants were statin-naive beneficiaries aged [&ge;]65 years hospitalized for myocardial infarction or cerebral infarction (primary diagnosis, length of stay [&ge;]3 days) and discharged home. Up to 14 nested daily trials (Trials 0-13) were constructed beginning on the discharge date, with eligibility, treatment assignment, and follow-up synchronized at each trial origin to eliminate immortal time. Primary analyses stacked all eligible trials; a pre-specified sensitivity analysis restricted inference to Trials 0 and 1, which achieved superior covariate balance (maximum standardized mean difference <0.1). Treatment Strategies: Initiation of atorvastatin (strategy A1) versus initiation of any other new outpatient medication (strategy A2). Strategy A0 (no new medication) was retained only to preserve sequential eligibility. Per-protocol effects were estimated after inverse-probability-of-treatment and inverse-probability-of-censoring weighting, with artificial censoring for treatment deviation (including a 30-day grace period) and death treated as a competing risk in Fine-Gray models. Main Outcomes and Measures: Previously detected signals and more granular, clinically coherent alternatives within the same outcome families (i.e., hemorrhagic events, cardiac valve disorders, musculoskeletal injuries, sensory symptoms, abnormal laboratory findings, and hyperglycemic events), defined by Clinical Classifications Software Refined categories plus independently validated Sentinel or published algorithms. Incident events required absence of relevant baseline codes. Confirmation required (1) within-outcome Benjamini-Hochberg q [&le;]0.05 with subdistribution hazard ratio (sHR) >1.0 and (2) both the median and 2.5th percentile of the bias-adjusted sHR remaining >1.0 across 5,000 Monte Carlo draws of probabilistic quantitative bias analysis (confounder-outcome risk ratio 1.25-3.00; prevalence difference 0.05-0.25). Absolute risks, risk differences, and numbers needed to harm (NNH) were reported. Stratified analyses examined time windows (1-30, 31-91, 92-182 days), age, sex, and race. Results: Of 70,130 eligible patients, 39,948 initiated atorvastatin and 19,182 initiated another new medication. After weighting, baseline covariates were closely balanced. Acute hemorrhagic cerebrovascular disease was confirmed overall (sHR 1.43, 95% CI 1.00-2.04; risk difference 0.5%; NNH 205) and more strongly in the first 30 days (sHR 2.20, 1.35-3.58); the association persisted in Trials 0 and 1 (sHR 1.50, 1.02-2.20). Related early intracranial hemorrhage signals were likewise confirmed. Nonrheumatic and unspecified valve disorders were confirmed in days 92-182 (sHR 1.48-1.58), as was cardiac valve intervention overall (sHR 1.74-1.83). Sprains, strains, and related composites were confirmed among men (sHR 1.66-1.94). General sensation/perception symptoms and dizziness were confirmed among non-White patients (sHR 1.40-1.43) but only in the unrestricted trial set. Acute hepatic failure was confirmed overall (sHR 1.61-1.72), and biliary tract disease among women (sHR 1.45-1.49). For every confirmed association the proportion of bias-adjusted draws remaining above the null was 1.00. Multiple prior signals, including prediabetes and acute posthemorrhagic anemia, failed the dual confirmation criteria. Conclusions: Sequential target-trial emulations with refined outcome definitions, within-outcome multiplicity control, restriction to optimally balanced early trials, and probabilistic quantitative bias analysis confirmed several ADE signals associated with atorvastatin initiation in older adults--most notably early hemorrhagic cerebrovascular events, cardiac valve disorders and interventions, musculoskeletal injuries in men, and selected hepatobiliary events--while attenuating others. Absolute excess risks were modest yet clinically relevant in a high-risk post-infarction population. These findings support a two-stage active pharmacovigilance paradigm (signal detection followed by rigorous confirmation) and justify heightened clinical vigilance for the confirmed events, while underscoring the need for external validation in independent populations and data sources.

8
Quality, consistency, and clinical safety of AI-generated versus clinician-written clinical notes: a multi-country paired simulation study

Bergman, H. I.; Liu, V.; Austin, B.; Ali, S.; Fiedler, M.; Sandiford, C.; Blanchard, R.; Casanovas, C. L.; Pedrazzini, G.; Markopouliotis, T.; Vermersch, F.

2026-08-21 health informatics 10.64898/2026.08.18.26360701 medRxiv
Top 0.2%
2.0%
Show abstract

Background Ambient AI documentation tools, known as scribes, are entering routine clinical practice at scale, but the evidence comparing the notes they produce against clinician-written notes is dominated by single-site, single-language studies that rely on human review to find errors, a method known to miss most documentation errors. Methods We conducted a paired simulation across five countries and languages (Cambridge/English, Barcelona/Spanish, Milan/Italian, Paris/French, Cologne/German; 385 paired consultations, 770 notes). From each actor-performed consultation, an AI scribe (Heidi) and a junior-to-middle-grade clinician independently produced a note. Notes were scored on the PDQI-9 by evaluators blinded to authorship. Documentation errors were identified by two methods of deliberately different sensitivity - clinician adjudication, and a calibrated automated reviewer externally validated against a blinded ten-clinician panel - then graded for clinical risk by a three-model panel. The co-primary outcomes were PDQI-9 total and Critical+High error burden, the latter reported under both detection arms. The analysis plan was registered before any pooling across sites. Results AI notes scored higher than clinician notes on the PDQI-9 (40.6 vs 35.6; difference +5.08, 95% CI 4.6-5.6; Cohen dz=0.55), consistently across all five sites (dz 0.41-0.75), and were less dispersed (5.7% of AI vs 27.8% of clinician notes fell below the study pre-specified low-score threshold (<32)). On the principal safety outcome - the paired probability that a note carried [&ge;]Critical+High error - clinician notes were affected more often under both detection arms: 61.0% versus 24.4% by the calibrated reviewer (relative risk 2.50, 95% CI 2.09-3.00) and 21.8% versus 6.2% by clinician adjudication (relative risk 3.50, 95% CI 2.32-5.27). The difference was largest for omissions. Unaided clinician review identified roughly 12% of the errors the calibrated reviewer retained, and a smaller fraction in AI notes than in clinician notes. Conclusions In this simulation, AI-generated notes scored higher on documentation quality, varied less, and carried fewer clinically significant errors than notes written on the same consultations by junior-to-middle-grade clinicians. The magnitude of the safety difference depends on the sensitivity of error detection, so we report both detection regimes and bound rather than point-estimate the absolute error rate. Extension to live practice, consultant-authored documentation, and notes as filed after clinician editing remains to be established.

9
Levodopa Administration Timing During Hospitalization: Associations With Intensive Care Unit Exposure and Documented Access Type

Gorenshtein, A.; Katson, M.; Adiniaev, Y.; Klang, E.; Daniel, O.

2026-08-19 neurology 10.64898/2026.08.17.26360596 medRxiv
Top 0.2%
1.9%
Show abstract

Background: Levodopa is time-critical in hospitalized Parkinson disease. Whether dosing fidelity depends on care setting or documented access status is unclear. Objectives: To quantify levodopa dosing fidelity, test ICU exposure with clustering-aware methods, and test whether documented access type is associated with delayed or omitted dosing. Methods: Retrospective cohort study using MIMIC-IV (2011-2022). Adults with Parkinson disease and [&ge;]1 scheduled levodopa dose contributed 1,665 admissions and 39,322 doses. ICU exposure was tested with a patient-clustered GEE model. Among ICU-exposed doses, access type (normal, tube feeding, parenteral nutrition, NPO) was modeled in one fully adjusted model and tested for specificity, restricted to the ICU, against an active-comparator medication (statins). Results: Of 39,322 doses, 79.8% were on time by the primary 60-minute definition; a symmetric {+/-}15-minute definition classified 68.8% as mistimed. ICU exposure was not associated with delayed or omitted dosing after clustering (patient-clustered OR, 0.87; 95% CI, 0.74-1.01). Among ICU-exposed doses, NPO was associated with delayed or omitted dosing (adjusted OR, 1.89; 95% CI, 1.36-2.62) and tube feeding with lower odds (adjusted OR, 0.62; 95% CI, 0.42-0.92; P < .001). The comparator medication showed a directionally consistent but inconclusive interaction (OR, 1.27-1.28; 92 patients). A route-order association was not observed among immediate-release formulations (OR, 0.72; 4 patients). Conclusions: ICU admission alone was not associated with dosing unreliability after clustering. Among ICU-exposed doses, access type, not a single pooled category, was associated with dosing reliability; a comparator-medication check, valid only in the ICU, was directionally consistent but inconclusive.

10
Trends in the Utilization of Breast, Cervical, and Colorectal Cancer Screening from 2010 to 2019 Among a Commercially Insured Population Using the MarketScan Commercial Claims Database

Sun, J.; Wat, R.; Frick, K. D.; Kong, X.; Liang, H.; Chow, C.; Shi, L.

2026-08-11 epidemiology 10.64898/2026.08.09.26360037 medRxiv
Top 0.2%
1.5%
Show abstract

Introduction: Breast, cervical, and colorectal cancer screening guidelines changed substantially between 2010 and 2019. We examined trends in the annual utilization of these screenings among commercially insured enrollees in the United States from 2010 to 2019 by age group, geographic region, and screening modality. Methods: We conducted a retrospective, serial cross-sectional analysis of the MarketScan Commercial Claims Database from 2010 through 2019, comprising approximately 141.2 million privately insured enrollees. Annual screening rates, defined as the proportion of eligible enrollees receiving a given test within each calendar year, were estimated for cervical, breast, and colorectal cancer using procedure codes, stratified by age group, screening modality, and geographic residence. These reflect annual utilization rather than up-to-date (guideline-concordant) screening. Temporal trends were evaluated using two-sided Poisson regression, and urban-rural disparities in 2019 were assessed using multivariate generalized estimating equations. Results: Cancer screening utilization remained stagnant or declined across all three cancer types over the study period. Among women aged 30-64 years, cervical cytology alone declined substantially from 28.2% in 2010 to 8.8% in 2019, while co-testing increased from 11.4% to 20.3%. Screening mammography among women aged 50-64 showed minimal change, remaining stable at 45.7% in 2010 and 45.8% in 2019. Colorectal cancer screening across enrollees aged <64 decreased modestly from 7.7% in 2010 to 6.5% in 2019, with a more pronounced decline among adults aged 45-49 years. Across all three cancer types, screening utilization was higher among urban residents than rural residents, with incidence rate ratios ranging from 1.02 to 1.05 in 2019. Conclusions: Utilization of cervical, breast, and colorectal cancer screening among commercially insured adults did not improve between 2010 and 2019. Persistent urban-rural disparities highlight ongoing gaps in preventive care delivery. Targeted interventions may help improve screening utilization, particularly in rural and underserved populations.

11
Practices and Perceptions Regarding Hyperkalemia in Clinical Practice: A Survey Study in Central America and the Dominican Republic

Ortiz, D. W.; Gonzalez, J.; Sanchez Polo, J. V.; Avellan, M.; Gonzalez, P.

2026-08-06 epidemiology 10.64898/2026.08.04.26359744 medRxiv
Top 0.3%
1.5%
Show abstract

Background: Hyperkalemia is a clinically relevant disorder across the cardiorenal continuum. In Central America and the Dominican Republic, there are no published systematic descriptions of real-world clinical practices or the degree of alignment of these practices with the most recent hyperkalemia management guidelines. Objective: To characterize physicians perceptions and therapeutic behaviors regarding hyperkalemia, including diagnostic thresholds, criteria for intervention and referral, management strategies, and access to potassium monitoring. Methods: A cross-sectional study was conducted using an online survey administered between April and June 2025 to physicians from multiple specialties across seven countries. Absolute and relative frequencies were calculated overall and stratified by specialty and country. Results: A total of 362 responses were collected. Participants were primarily from Costa Rica (32.3%), Honduras (27.9%), and Guatemala (21.0%). 37.8% of respondents reported hyperkalemia in 10% to 30% of their patients, with the most reported diagnostic threshold being serum potassium 5.5 mEq/L. Outpatient intervention was most frequently initiated at 5.5 mEq/L (55.2%), while referral to the emergency department was reported at a potassium level of 6.0 mEq/L (35.6%). Regarding management strategies, 67.0% favored an electrocardiogram prior to deciding on intervention; 93.0% reported reduction or discontinuation of drug causing hiperkalemia; and 74.0% prescribed therapies increasing potassium excretion. Access to potassium monitoring differed substantially by setting reported as 55.5% in the public versus 90.3% in the private sector. Among cardiologists, frequently used strategies for hyperkalemia in heart failure were reduction or discontinuation of mineralocorticoid receptor antagonists and increased use of loop diuretics. Nephrologists favored strict dietary modifications, loop diuretics, and the use of cation-exchange resins. Conclusions: Substantial heterogeneity was observed in hyperkalemia definitions, action thresholds, and referral criteria, along with frequent modification of renin-angiotensin-aldosterone inhibitors, and reduced access to potassium monitoring in the public sector.

12
Social Determinants of Health and GLP-1 RA Use Among Patients with Type 2 Diabetes and Stage 2 CKM: Insights from NHANES 2005?2020

Jian, Q.; Segal, M. S.; Shao, H.; Singh-Ospina, N.; Jiao, T.

2026-08-21 epidemiology 10.64898/2026.08.18.26360763 medRxiv
Top 0.3%
1.5%
Show abstract

Background Cardiovascular-Kidney-Metabolic (CKM) syndrome encompasses interconnected conditions such as type 2 diabetes (T2D), hypertension, hypertriglyceridemia, metabolic syndrome (MetS), and chronic kidney disease (CKD). As CKM progresses, cardiorenal risks increase. Although Glucagon-like peptide-1 receptor agonists (GLP-1 RA) have demonstrated cardiorenal and cardiometabolic benefits, offering an opportunity to slow CKM progression, their use may vary across social determinants of health (SDoH) and stage 2 CKM subgroups. Objective To evaluate the influence of SDoH on access to GLP-1 RA among patients with T2D and other stage 2 CKM conditions. Methods This cross-sectional study used data from the U.S. National Health and Nutrition Examination Survey (NHANES), 2005?2020. Adults aged [&ge;]30 years with T2D and/or other stage 2 CKM conditions were included. Weighted descriptive analysis, multivariable logistic regression and LASSO were applied to assess associations between SDoH and GLP-1 RA use. Results Among 4,520 participants (representing approximately 84.0 million U.S. adults), weighted mean age was 61.4 years, 48.9% were female, and 61.5% were non-Hispanic White. Among participants with T2D, GLP-1 RA use was higher among individuals with higher education (3.39% vs 1.43%), private insurance (3.00% vs 0.58%), and higher income (4.70% vs 1.87%), while no use was observed among those without routine places for care. In adjusted analyses, individuals with lower income, less than high school education, lack of insurance, and being unmarried had 64%, 51%, 81%, and 40% lower likelihood of GLP-1 RA use, respectively. LASSO identified income, education, insurance, and access to care as predictors. Lower income, lower educational attainment, and lack of insurance were associated with 48%, 34%, and 79% lower likelihood of GLP-1 RA use, respectively, adjusting for age, sex, and race/ethnicity. Conclusion SDoH-driven disparities limit GLP-1 RA access. Expanding GLP-1 RA access by addressing socioeconomic barriers is critical to slowing CKM progression, reducing cardiovascular risk, and mitigating health disparities.

13
Patient and Clinician Perspectives on Centralized Cascade Screening for Familial Hypercholesterolemia in the United States: A Qualitative Implementation Study

Roberts, M. C.; Jones, L. K.; Brown, A.; Carda-Auten, J.; Cuchel, M.; Hilton, A. R.; Khera, A.; Rothstein, M.; Soe, K.; Sullivan, A.; Tricou, E.; Vu, M. B.; Weintraub, W. S.; Ahmad, Z.

2026-08-19 cardiovascular medicine 10.64898/2026.08.18.26359725 medRxiv
Top 0.3%
1.5%
Show abstract

Objective: To identify patient- and clinician-reported barriers, facilitators, and design requirements for a centralized cascade-screening program for familial hypercholesterolemia (FH) in the United States. Methods: From June through November 2023, we conducted individual telephone interviews with 20 patients with FH and 10 clinicians recruited from UT Southwestern Medical Center, Parkland Health, the North Texas Veterans Affairs, and other clinical settings. Interview guides were informed by the Consolidated Framework for Implementation Research. Transcripts were coded in Dedoose using a piloted codebook, with discrepancies and emergent themes resolved through consensus. An advisory panel then helped translate interview findings into program design requirements and implementation strategies. Results: Five themes characterized barriers and facilitators to centralized cascade screening: (1) health-system access and fragmentation, including screening and treatment costs, transportation, and cross-system coordination; (2) privacy and trust, including concerns about genetic information and unsolicited outreach; (3) family relationships and practical burden, including competing demands, language barriers, limited contact, fear, and denial; (4) clinician capacity and workflow, including limited time, knowledge, and genetic-counseling capacity; and (5) communication and care continuity. Participants recommended proband pre-notification of relatives, culturally and linguistically responsive materials, secure data exchange, standardized scripts, flexible testing pathways, and centralized coordination. These findings informed a program model incorporating a secure pedigree platform, educational and communication resources, testing coordination, and linkage to follow-up care. Conclusions: Patients and clinicians identified multilevel determinants that a centralized FH cascade-screening program must address. The findings support specific design requirements but do not establish program feasibility or effectiveness, which require prospective evaluation.

14
The Current State of Timely Results Reporting Among Hypertension Trials on ClinicalTrials.gov: A Cross-Sectional Meta-Research Analysis

Harris, W. T.; Bragg, P.; Kocour, L.; Livsey, T.; Langerman, R.; Calvert, N.; Lackey, M.; Nguyen, A.; Ford, A.; Vassar, M.

2026-08-07 cardiovascular medicine 10.64898/2026.08.05.26359829 medRxiv
Top 0.3%
1.3%
Show abstract

Objectives: To characterize how completely and promptly summary results are reported for registered hypertension trials on ClinicalTrials.gov, and whether reporting correlates with the observable obligation to report. Methods: Cross-sectional analysis of completed or terminated interventional trials for hypertension, retrieved through the ClinicalTrials.gov API version 2. Trials required a primary completion date of type ACTUAL at least 12 months before extraction. Reporting was timed from primary completion to first results submission and classified as timely at 365 days or fewer. Applicability was approximated requiring interventional design, phase 2 or later, a United States site, and an FDA-regulated drug or device, assigned flag-confirmed or inferred. Proportions are reported with Wilson 95% confidence intervals, time to reporting by Kaplan-Meier, and adjusted associations by logistic regression clustered on lead sponsor. Results: Of 5,851 trials, 5,396 were due to report. Timely reporting was 9.1% (95% CI 8.3-9.9) and any-time reporting 28.8% (95% CI 27.6-30.0). Reporting was graded by applicability, with flag-confirmed trials reporting timely at 36.9% (95% CI 31.6-42.5) and non-applicable trials at 6.3% (95% CI 5.6-7.1). A United States site carried the strongest adjusted association with timely reporting (OR 4.03, 95% CI 2.99-5.42). Among unreported trials, 7.8% had a sponsor-tagged publication and 36.4% under a broader definition. Conclusion: Prompt registry reporting of hypertension trial results remains uncommon, and reporting is most closely associated with the observable obligation to report.

15
Neuro-Adverse Events Associated with GLP-1 Receptor Agonists: A Study Based on the FAERS Database and External Validation Using NHANES Database

Bai, L.; Liu, Y.; Tongye, H.

2026-08-06 health economics 10.64898/2026.08.04.26359670 medRxiv
Top 0.3%
1.2%
Show abstract

Background Glucagon-like peptide-1 receptor agonists (GLP-1RAs) are widely prescribed for type 2 diabetes and obesity, yet their neuropsychiatric safety profile remains incompletely characterized. We aimed to systematically evaluate neuro-adverse event (AE) signals for six GLP-1RAs and to validate key findings using population-based data. Methods We conducted disproportionality analysis of FAERS data for semaglutide, liraglutide, dulaglutide, tirzepatide, exenatide, and lixisenatide. RORs were calculated for 93 predefined neuro-AE MedDRA PTs across 11 neurological categories. External validation used NHANES 2013-2018 (n=17,057; 70 GLP-1RA users) with survey-weighted regression. Results We identified 41 significant neuro-AE signals. Semaglutide showed the strongest neuromuscular signal, muscle atrophy (ROR 3.94; 95%CI 3.42-4.54), corroborated by tirzepatide (ROR 2.35; 95%CI 2.04-2.71). Exenatide generated the highest psychiatric signal: nervousness (ROR 4.03; 95%CI 3.70-4.40). NHANES confirmed higher depression odds (OR 2.05; 95%CI 1.32-3.19; P=0.001) and reduced sleep hours (beta -0.35; P=0.033). Conclusions GLP-1RAs carry multiple neuropsychiatric safety signals, including muscle atrophy as a potential class effect and depression risk corroborated by population-level data. These findings support heightened clinical monitoring.

16
Measuring implementation of clinical guidelines through the COVID-19 pandemic, using linked national health records: a national study of type 2 diabetes in England

Biglarbeigi, P.; Dale, C.; Lambarth, A.; Mason, A.; Takher, R.; Ballabio, G.; Minshull, J.; Mamas, M. A.; Tomlinson, C.; Rowark, S.; Rayman, G.; Pearson, E. R.; Khunti, K.; Sattar, N.; Sofat, R.

2026-08-10 health policy 10.64898/2026.08.07.26359950 medRxiv
Top 0.3%
1.2%
Show abstract

Objectives: To examine the conformance to type 2 diabetes NICE guidelines across cardiovascular risk strata; and to quantify geographical variation in treatment pathways following the COVID-19 pandemic, encompassing guideline changes. Design: We carried out a retrospective observational study using linked electronic health records across England. Process mining, a data driven method that can reconstruct clinical treatment pathways, was applied to map 12-month treatment trajectories after treatment initiation. Conformance with NICE NG28 (2022) was quantified using a structural similarity index. Further, behavioural and entropy-based similarity (capturing treatment variability and complexity) measures were used to assess sequencing and heterogeneity of treatment. Setting: Primary and secondary care in England datasets within the National Health Service England Secure Data Environment (NHSE SDE), analysed first at national level and then across 42 Integrated Care Boards (ICBs) which are the devolved health care geographical delivery regions in England. Participants: 822,650 individuals with newly diagnosed T2DM between 1-February-2022 and 1-November-2025, stratified into low cardiovascular risk (LR-C; QRISK3<10), high risk (HR-C; QRISK3>=10 or receiving statins/blood pressure lowering treatment), and established cardiovascular disease (eCVD-C). Participants were followed for 12 months after first dispensed glucose lowering therapy. Main outcome measure: First line therapy, treatment intensification and switching within 12 months; change in glycated haemoglobin (HbA1c); quantified conformance to NICE recommended pathways; and regional variation in broader similarity measures. Results: Metformin monotherapy was the dominant initiation strategy in LR-C and HR-C cohorts (92.4% and 90.2%, respectively), whereas eCVD-C showed lower uptake of metformin (68.9%) and higher uptake of SGLT2 inhibitors (26.3%). Intensification from metformin to combination therapy was infrequent across all cohorts (<1%), although HR-C demonstrated the highest treatment transitions and switching behaviour. Dispensed SGLT2 inhibitor use was nearly threefold higher in eCVD-C (26.7%) than in LR-C (9.0%) or HR-C (10.7%). Overall, conformance to NICE-recommended pathways remained modest nationally, particularly in LR-C and HR-C. Across 42 ICBs, substantial regional heterogeneity in treatment pathways and guideline conformance was observed, with conformance ranging from 0.29 to 1.00 in LR-C pathways, 0.40 to 1.00 in HR-C pathways, and 0.54 to 0.92 in eCVD pathways. Conclusion: National T2DM treatment pathways for post-pandemic showed higher alignment to NICE guidelines in eCVD-C compared to the other risk groups, with ongoing gaps and large regional variations in other risk groups. Process mining offers a scalable approach to monitor implementation of guideline recommended care that could support learning health systems. Using T2DM during and post COVID-19 pandemic as a case study, this work demonstrates how these methods can assess the use of existing and innovative therapies, identify gaps and guide future adoption to ensure recommended treatments reach the right patient groups.

17
TrialCode Agent: LLM-Assisted Clinical Code-Set Construction for Trial Emulation

Habibdoust, A.; Sajjad, A.; Hernandez, D.; Patel, K.; Song, X.

2026-08-23 health informatics 10.64898/2026.08.20.26360962 medRxiv
Top 0.4%
1.1%
Show abstract

Objective Translating free-text clinical trial criteria into computable code sets is a valuable standardization practice that is necessary for producing reproducible real-world evidence studies but requires standardized interpretation across multiple clinical vocabularies. Methods We developed TrialCode Agent, a hybrid-large language model (LLM)-terminology verification agent that generates, formats, verifies, and expands candidate codes from free-text clinical criteria. The system supports ICD-9-CM diagnoses and procedures, ICD-10-CM, ICD-10-PCS, LOINC, and RxNorm medication concepts. We compared Baseline, Hybrid biomedical retrieval-augmented generation (RAG), and terminology-guided Family expansion pipelines using Claude, GPT Qwen, and MedGemma on 40 criteria from 11 trial groups. Performance was evaluated against expert-built reference code sets using exact-code precision, recall, and F1. Results The optimal pipeline varied by model. Claude with Baseline achieved the highest performance (precision 0.755, recall 0.619, F1 0.680), followed by GPT-5.5 with Baseline (precision 0.569, recall 0.658, F1 0.610), Qwen with Hybrid biomedical RAG (precision 0.656, recall 0.470, F1 0.548), and MedGemma with Family expansion (precision 0.487, recall 0.316, F1 0.383). Hybrid biomedical RAG improved aggregate F1 only for Qwen but increased GPT-5.5 RxNorm F1 from 0.320 to 0.909. Macro-averaged results showed criterion-level gains despite lower micro-averaged aggregate performance. Family expansion increased recall across models but generally reduced precision. In staged verifier ablation, micro-F1 increased from 0.254 before verification to 0.505 after final verification and expansion. Existence/vocabulary checking removed 2,594 false-positive codes, and acceptance filtering removed 952 additional false-positive codes before controlled expansion. Conclusions Combining LLM-based clinical interpretation with deterministic terminology verification produces auditable, database-ready code sets, but retrieval and broad family expansion do not consistently improve exact-code performance. Retrieval was particularly useful for RxNorm mapping, whereas overly broad or incomplete candidate generation remained the main source of error. Deterministic verification improves code validity and query readiness but cannot replace accurate clinical interpretation.

18
The Heartbeat Study: Feasibility and Advertisement Costs of Implementing a Digital Strategy to Enhance Diversity in the LIBREXIA-AF Clinical Trial

Hussain, T.; Wang, Y.; Chen, Y. Q.; Olson, G.; Panitch, B.; Clemins, K.; Elkarra, N.; Lhamo, K.; Odenwald, N.; Hufner, D.; Jain, S.; Quall, M.; Anderson, C.; Perez, M. V.

2026-08-28 cardiovascular medicine 10.64898/2026.08.24.26361277 medRxiv
Top 0.4%
1.1%
Show abstract

Background: Recruitment of diverse participants remains a challenge in cardiovascular clinical trials. Little is known about how recruitment efficiency and advertising costs with web-based tools vary across US communities. We evaluated an online recruitment platform and examined the cost of acquiring both all-comers and diverse participants in relation to community-level income. Methods: The Heartbeat Study evaluated a digital recruitment strategy to identify US participants for the ongoing Phase 3 LIBREXIA-AF trial. Online advertisements directed individuals with atrial fibrillation to a pre-screening website, where demographic and health data were collected. Advertising impressions, clicks, and costs were recorded. Participant ZIP codes were linked to Core Based Statistical Areas (CBSAs) and CBSA-level income. We measured recruits from underrepresented groups (women, African Americans, Latinos) completing online registration per $100,000 in advertising expenses. Click-weighted linear regression evaluated associations between CBSA income and advertising efficiency. Results: A total of 1,406 recruits completed online registration, with 1,319 participants from 260 CBSAs included in the geographic analysis. Participants were 73 years old on average; 547 (41.5%) were women, 59 (4.5%) African American, and 44 (3.3%) Latino. A total of $163,949.13 was spent on 82,681,711 impressions and 454,750 clicks. Recruits per $100,000 in advertising spend were 334 for women, 36 for African Americans, and 27 for Latinos. CBSA-level income was modestly inversely associated with cost per impression (R2=0.058; p<0.001) and cost per click (R2=0.038; p=0.005), but not recruitment yield for African Americans (p=0.99), Latinos (p=0.37), or women (p=0.21) (R2 range, 0.000-0.13). Conclusions: In this national analysis, online advertising enabled broad engagement across diverse US communities, but income was not associated with recruitment yield among women, African American, or Latino participants. Minority representation remained limited, suggesting digital recruitment alone may be insufficient to improve trial diversity. Targeted, culturally and linguistically tailored strategies may be needed to enhance diverse recruitment.

19
Machine Learning-Supported Efficient VTE Risk Assessment using Routinely Collected Electronic Health Record Data

Li, Z.; Yagis, E.; Riad, A.; Windrath-Carr, O.; Arribas, M.; Sodiq, T.; Goldsmith, K.; Glampson, B.; Flott, K.; Haji, G.; Khan, Z.; Baker, C.; Mayer, E. K.

2026-08-19 health informatics 10.64898/2026.08.18.26360687 medRxiv
Top 0.4%
1.1%
Show abstract

Venous thromboembolism (VTE) is a leading cause of preventable inpatient mortality, while the real-world performance of mandated risk assessment and the potential for automating using electronic health record (EHR) data remain unclear. We analysed 577,904 admissions and 726,896 VTE assessment forms across five NHS hospitals between 2015 and 2025 to evaluate assessment completion, concordance with structured EHR data, clinical validity, and feasibility of EHR-based automation assisted by machine learning. Overall completion was high (96.7%), and timely completion improved from 47.4% in 2015 to 90.5% in 2024. Agreement between forms and EHR data was good for common risk factors, but low-prevalence variables were often under-documented in the forms. Despite these discrepancies, form-derived thrombosis risk was associated with increased VTE incidence (OR 3.31, 95% CI 2.81-3.90). Machine learning models using first-14-hour EHR data achieved discrimination comparable to clinician-recorded variables (AUROC 0.709 vs 0.704), supporting real-time EHR-integrated assessment pre-population and decision support.

20
A Pragmatic Randomized Trial of an EHR-Integrated Generative AI Chart Summarization Tool for Ambulatory Clinicians

Chin, A. T.; Zhu, N.; Vangala, S.; Woo, H.; Wisk, L. E.; Kingsley, T.; Mafi, J. N.; Lukac, P. J.

2026-08-31 health informatics 10.64898/2026.08.26.26361496 medRxiv
Top 0.4%
1.1%
Show abstract

BACKGROUND Generative AI (genAI) chart summarization tools embedded in electronic health records (EHRs) are being rapidly deployed across U.S. health systems. Although these tools represent a promising solution to alleviate cognitive burdens, their effects have not been examined in randomized-clinical trials (RCTs). METHODS In this pragmatic RCT at a single academic health system, 284 outpatient clinicians across forty-two specialties were assigned 1:1 to Epic's outpatient chart summarization tool or a usual-care control arm over 90 days, from February 23 to May 23, 2026. The primary outcome was physician task load (PTL) adapted for pre-charting. Prespecified exploratory outcomes included additional validated psychometrics as well as usability, safety, and time-based measures. Descriptive statistics included interaction and usage of the tool. RESULTS Of 74,474 AI chart summaries generated, 14.2% were interacted with by a clinician; the proportion of generated summaries interacted with declined from 21.5% in month 1 to 10.5% in month 3, and the proportion of clinicians using the tool at least once per month declined from 88.7% to 66.2%. The adjusted between-arm difference in PTL at follow-up favored the intervention arm (scale 0-400; -27.4; 95% CI, -49.4 to -5.3; P=0.02). Among the Professional Fulfillment Index (PFI; scale 0-4, lower=better) psychometrics, overall burnout (-0.20; 95% CI, -0.38 to -0.01) and work exhaustion (-0.24; 95% CI, -0.47 to -0.02) were lower in the intervention arm, with little difference in overall professional fulfillment (+0.04; 95% CI, -0.16 to 0.25). Charting time per encounter showed no significant between-arm difference during steady state (-1.2 seconds; 95% CI, -19.0 to 16.6). The net promoter score was -22, indicating that on average, clinicians did not recommend the tool. Among free-text respondents, 57.1% reported at least one concern, most commonly tool limitations or inaccurate information. No adverse patient safety events or near-misses were reported. CONCLUSION An EHR-integrated AI chart summarization tool modestly reduced physician task load and was associated with lower burnout, without time savings and against declining engagement. Sustained usage and oversight of reported inaccuracies remain open challenges.